Skip to content

fix(ci): pin runpodctl v2.9.0 and drop the deprecated config step (RUNPOD_API_KEY auth) - #470

Open
fusheng-ji wants to merge 3 commits into
RL-Align:mainfrom
fusheng-ji:fix/ci-runpodctl-auth
Open

fusheng-ji wants to merge 3 commits into
RL-Align:mainfrom
fusheng-ji:fix/ci-runpodctl-auth

Conversation

@fusheng-ji

@fusheng-ji fusheng-ji commented Oct 4, 2026 •

Copy link
Copy Markdown

Latest Status [2026-10-04]

2026-10-04: rebased onto main (6d8b4bc) and signed off (DCO): one commit, a1736d1, with an
identical patch (git patch-id). The results below were measured before the rebase.

Ready for review. One commit, three workflow files, no script changes.

Line numbers refer to this branch (a1736d1). The workflow files shift by a few lines
compared with main; ci/run_gpu_ci.sh is unchanged.

Summary

TL;DR: ws1-gtest-gpu and ws1-chain-gpu fail at Configure runpodctl before any
test runs. ws1-gtest-gpu has never succeeded. This PR deletes that step and pins
runpodctl v2.9.0 with a sha256 check. runpodctl v2 already reads RUNPOD_API_KEY,
and the step that runs ci/run_gpu_ci.sh already exports it.

  • Cause: step order. The workflows install releases/latest (v2.14.0 since
    2026-09-10) and then run runpodctl config --apiKey … as the first runpodctl call on
    a fresh runner. That fails because no config file exists yet. gpu-ci.yml gets past
    the same command only because it runs runpodctl version first, which creates the file.
  • Fix, applied identically to all three workflows:
    1. Pin v2.9.0. It is the version a previous CI run
      (31455127310)
      exercised end to end: pod create, pod-id parsing, pod get polling, SSH,
      pod remove. The download is verified with sha256sum -c, and runpodctl version
      is printed into the log.
    2. Delete Configure runpodctl. Every later runpodctl call is inside
      ci/run_gpu_ci.sh, and every step that runs it sets RUNPOD_API_KEY in its env:.
      That includes the pod remove in its EXIT trap, which runs in the same shell.
  • Side effect removed. When config succeeded (in gpu-ci, which runs version
    first), it also generated an SSH key pair and uploaded it to the RunPod account. The script never used it: ssh/scp pass no -i
    and use the key written by Setup SSH key.

Files

file status
.github/workflows/ws1-gtest-gpu.yml pinned install (:75-88); Configure runpodctl removed; key at :98, script at :106
.github/workflows/ws1-chain-gpu.yml same; key at :93, script at :102
.github/workflows/gpu-ci.yml same; key at :80, script at :86

git diff --stat upstream/main (1968a87): 3 files, 29 insertions, 12 deletions.

Test

# The failure on main (read-only API calls; gh 2.x)
gh run view -R RL-Align/RL-Kernel 36835086246 --log-failed
# Runs by conclusion; as of 2026-10-01: 69 failure, 40 skipped, 13 action_required, 0 success
gh api --paginate 'repos/RL-Align/RL-Kernel/actions/workflows/ws1-gtest-gpu.yml/runs?per_page=100' \
    -q '.workflow_runs[] | .conclusion // .status' | sort | uniq -c

# The pinned binary
curl -fsSLO https://github.com/runpod/runpodctl/releases/download/v2.9.0/runpodctl-linux-amd64
curl -fsSLO https://github.com/runpod/runpodctl/releases/download/v2.9.0/checksums_2.9.0_sha256.txt
grep ' runpodctl-linux-amd64$' checksums_2.9.0_sha256.txt | sha256sum -c -
chmod +x runpodctl-linux-amd64 && ./runpodctl-linux-amd64 version
./runpodctl-linux-amd64 pod create --help      # lists every flag run_gpu_ci.sh passes

# v2.9.0 reads RUNPOD_API_KEY: run with no network at all, so no request can leave.
# unshare -rn needs unprivileged user namespaces, which some distributions restrict.
H=$(mktemp -d)
unshare -rn env -i PATH=/usr/bin:/bin HOME="$H" bash <<'SH'
ip -o link                                                # only lo, and it is down
curl -sS -m5 https://1.1.1.1 || true                      # expect a connection failure
./runpodctl-linux-amd64 version
./runpodctl-linux-amd64 pod list                          # expect no_credentials
RUNPOD_API_KEY=dummy ./runpodctl-linux-amd64 pod list     # expect network_error
cat "$HOME/.runpod/config.toml"                           # apikey stays empty
SH

python -c 'import sys, yaml; [yaml.safe_load(open(f)) for f in sys.argv[1:]]' \
    .github/workflows/{ws1-gtest-gpu,ws1-chain-gpu,gpu-ci}.yml
actionlint .github/workflows/{ws1-gtest-gpu,ws1-chain-gpu,gpu-ci}.yml
pre-commit run --files .github/workflows/{ws1-gtest-gpu,ws1-chain-gpu,gpu-ci}.yml

Test results

Run locally against the pinned binary in a scratch directory. No pod was created, and
no request carrying a credential was sent to RunPod.

check result
failure on main run 36835086246: 'runpodctl config' is deprecated, then error saving config: Config File ".runpod.yaml" Not Found, exit 1
ws1-gtest-gpu history as of 2026-10-01, 0 successes in 122 runs (69 failed, 40 skipped, 13 action_required); its first runs (2026-08-19) already failed at this step
checksum runpodctl-linux-amd64: OK; sha256 06e6f549…a64148 also matches GitHub's asset digest
version runpodctl 2.9.0-c094cac, the same string run 31455127310 printed
flags all 7 pod create flags present; pod get -o json; pod remove is an alias of pod delete
reads RUNPOD_API_KEY no key: no_credentials; dummy key: network_error at DNS; config.toml stays apikey = ''; v2.14.0 identical
install step replay each workflow's install run: block, executed in a scratch directory (script not attached): all three print OK and the version; a wrong checksum prints FAILED and exits 1
YAML / lint yaml.safe_load passes; no Configure runpodctl left; actionlint 1.7.12 clean (shellcheck was not installed, so actionlint's checks of the run: shell code did not run); pre-commit passes
Local reproduction of the root cause (v2.14.0, empty HOME)
$ runpodctl config --apiKey dummy                       # config as the first call
{"error":"error saving config: Config File \".runpod.yaml\" Not Found in \"[<HOME>/.runpod <HOME>]\"","code":"cli_error"}

$ runpodctl version                                     # creates <HOME>/.runpod/config.toml
$ runpodctl config --apiKey dummy
Configuration saved to file: <HOME>/.runpod/config.toml

gpu-ci run 34585491929
(2026-09-11) shows the same thing in CI: it ran version first and then got past
Configure runpodctl. A later step (pod create) failed in that run for an unrelated
reason.

Notes

Risks this PR cannot verify
  1. v2.9.0's output against today's API. run_gpu_ci.sh parses runpodctl output
    with grep:

    • "no longer any instances available" (:60, :76);
    • the pod-id regexes (:82, :84);
    • the IP and port greps (:100, :101);
    • "not found" (:37).

    Run 31455127310 (2026-08-11, this exact binary) exercised :60/:76, :82,
    :100/:101 and the pod remove call. The :84 fallback (only reached when :82
    does not match) and the :37 branch were not exercised.

  2. Existing bug, not changed here. The pod-id fallback regex at :84 can take a
    JSON error code such as unauthorized as the pod id, and then poll it for up to
    10 minutes. A follow-up could check for a top-level "error" key first.

How this PR can be exercised before merge
  • From a fork, ws1-gtest-gpu.yml:59 and ws1-chain-gpu.yml:46 skip the GPU jobs at PR
    time.
  • gpu-ci.yml runs on pull_request_target (:4) and checks out its orchestrator from
    the base commit (:40). A needs-gpu-ci label would therefore run main's
    orchestrator, and its setup already passed before this PR.
  • The first real run is the push to main after merge: ws1-chain-gpu.yml has no
    paths filter, and ws1-gtest-gpu.yml's push paths include the workflow file
    (:47). A maintainer can also run workflow_dispatch on ws1-gtest-gpu.yml (:48)
    from a branch in RL-Align/RL-Kernel.
  • The SSH keys that past config runs uploaded may be worth pruning from the account.
    The generated public key carries the comment runpodctl-ssh-key.

Summary by CodeRabbit

  • Chores
    • GPU build and test workflows now use a specific version of the command-line tool and verify its checksum before installation.
    • The workflows no longer install the moving “latest” release or run a separate API-key configuration step.

ws1-gtest-gpu and ws1-chain-gpu fail at "Configure runpodctl" before any
test runs. They install releases/latest (v2.14.0 since 2026-09-10), whose
deprecated `config --apiKey` aborts with
`Config File ".runpod.yaml" Not Found` when it is the first runpodctl
call on a fresh HOME. It only succeeds once an earlier call has created
~/.runpod/config.toml, which is why gpu-ci, whose install step happens to
run `runpodctl version` first, got past the same command.

Drop the config step in all three workflows. runpodctl v2 reads
RUNPOD_API_KEY from the environment, and every step that runs
ci/run_gpu_ci.sh already exports it, including the cleanup trap that
calls `pod remove`. Verified for v2.9.0 inside `unshare -rn`, from a
fresh HOME after `runpodctl version`: without the variable `pod list`
fails locally with no_credentials, with RUNPOD_API_KEY=dummy it gets
past the credential check to a network_error, and config.toml keeps
`apikey = ''` throughout. v2.14.0 behaves the same.

Pin the download to v2.9.0, the version gpu-ci run 31455127310 exercised
end to end (pod create with sold-out fallback, pod id parsing, pod get
polling, SSH, cleanup), and verify it with sha256sum -c against the
runpodctl-linux-amd64 line of checksums_2.9.0_sha256.txt
(06e6f54957db79d5cd9f1909a7f1d365076826751ba2f5df65d75dde43a64148),
which matches GitHub's asset digest.

Signed-off-by: Wenbo Ji <36562829+fusheng-ji@users.noreply.github.com>
@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Important

Review skipped

Review was skipped as selected files did not have any reviewable changes.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 36051f2f-dcca-4b09-817e-c8d23d31db6d
📥 Commits

Reviewing files that changed from the base of the PR and between bcd49df and 116b9a1.

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 6cb4cc73-21e0-421e-87e0-49aa571747b3
📥 Commits

Reviewing files that changed from the base of the PR and between 11cac8c and bcd49df.

📒 Files selected for processing (3)
  • .github/workflows/gpu-ci.yml
  • .github/workflows/ws1-chain-gpu.yml
  • .github/workflows/ws1-gtest-gpu.yml

Included review availability: This review used your included allowance. Your plan provides up to 2 included reviews per hour; 1 remain after this review.


📝 Walkthrough

Walkthrough

Three GPU workflows now download runpodctl v2.9.0 and verify its SHA-256 checksum before installation. They remove separate API-key configuration steps. Two workflows print the installed version.

Changes

GPU workflow runpodctl installation

Layer / File(s) Summary
Pin and verify runpodctl
.github/workflows/gpu-ci.yml, .github/workflows/ws1-chain-gpu.yml, .github/workflows/ws1-gtest-gpu.yml
All three workflows download runpodctl v2.9.0 and verify its SHA-256 checksum. The chain and gtest workflows also print the installed version. Each workflow removes its separate API-key configuration step.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Merge Risk: ⚪ Minimal · up to bcd49

The workflows install the verified v2.9.0 binary and provide its required API key to the GPU script. No concrete regression from these changes remains; the PR is mergeable after normal checks.

Architecture Summary

Architecture risk: 🔵 Low · up to bcd49

The changed surface does not map to a changed system, dependency edge, entrypoint, or external dependency.

Changed systems: None identified.

Architecture concerns
No architecture-level concerns identified.

Review details

Before / after behavior

  • observed — Modified behavior in .github/workflows/gpu-ci.yml: The installation step now sets a fixed runpodctl version and SHA-256, downloads that version, and verifies the binary before installing it. It replaces the download from the moving latest release. The separate API-key configuration step was removed.
  • observed — Modified behavior in .github/workflows/ws1-chain-gpu.yml: The install step now downloads the v2.9.0 binary, checks it against the pinned SHA-256 value before making it executable and installing it, then runs runpodctl version. The previous download from the moving latest release was removed, as was the subsequent Configure runpodctl step that configured the API key.
  • observed — Modified behavior in .github/workflows/ws1-gtest-gpu.yml: The install step now sets a fixed runpodctl version and checksum, downloads that version, fails checksum verification if the binary does not match, and prints the installed version. It removes the floating latest download and the separate API-key configuration step.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main changes: pinning runpodctl v2.9.0 and removing the deprecated config step.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@fusheng-ji fusheng-ji changed the title fix(ci): pin runpodctl v2.9.0 and authenticate via RUNPOD_API_KEY fix(ci): pin runpodctl v2.9.0 and drop the deprecated config step (RUNPOD_API_KEY auth) Oct 4, 2026
@Flink-ddd Flink-ddd added the type: ci-cd Modify GitHub Actions, automated tests, and packaging/deployment tasks. label Oct 4, 2026
@fusheng-ji

Copy link
Copy Markdown
Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
⚠️ Action not completed

Already reviewed the last commit. Use @coderabbitai full review to rerun a review of the entire changeset.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@Flink-ddd Flink-ddd left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you run this on RunPod and share logs showing that pod creation, SSH connection, and pod deletion all work?

@fusheng-ji

fusheng-ji commented Oct 6, 2026 •

Copy link
Copy Markdown
Author

Thanks for the review. I don't have a RunPod account, and the GPU jobs can't run from a fork, so this needs a maintainer run with the org key. Here is what is available now and the run that would close the gap.

1. Existing evidence with the pinned binary. Run 31455127310 (GPU CI, 2026-08-11, success) used runpodctl 2.9.0-c094cac, the exact binary this PR pins, and completed all three stages in both jobs:

  • create: 2x RTX A4000 sold out → fallback 1x NVIDIA A40 → Successfully rented pod
  • SSH: Pod infrastructure is 100% READY! → remote suite ran (nvidia-smi output) → Remote execution finished with exit code = 0
  • delete: AUTOMATIC CLEANUP → {"deleted": true, ...}

What that run does not prove. In that run Configure runpodctl also wrote config.toml, so it does not show auth working from RUNPOD_API_KEY alone, which is what this PR depends on. Locally, without network access, v2.9.0 reports no_credentials with no key and attempts the request when RUNPOD_API_KEY is set. That has not been checked against the real API.

2. The failure this PR fixes is widespread. Across the 111 forks there are 138 runs of the GPU workflows: 108 failed, 30 skipped, 0 succeeded. A sample of 25 forks shows every failed job stopping at Configure runpodctl. Upstream, ws1-gtest-gpu has never succeeded.

3. Request. Could a maintainer push this PR's head (116b9a1) to a branch in RL-Align and run workflow_dispatch on ws1-gtest-gpu.yml from that branch? That run uses this PR's workflow file and the org RUNPOD_API_KEY. I'll go through the log and post the create, SSH and delete lines here.

For reference, a maintainer can dispatch it with:

git fetch https://github.com/fusheng-ji/RL-Kernel.git fix/ci-runpodctl-auth
git push upstream FETCH_HEAD:refs/heads/ci/pr-470-runpod-check
gh workflow run ws1-gtest-gpu.yml -R RL-Align/RL-Kernel --ref ci/pr-470-runpod-check```

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type: ci-cd Modify GitHub Actions, automated tests, and packaging/deployment tasks.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants